Papers with preference prediction

    1 papers
    ReflectRM: Boosting Generative Reward Models via Self-Reflection within a Unified Judgment Framework (2026.acl-long)

    Copied to clipboard

    Challenge: Existing methods for generating reward models focus on outcome-level supervision, neglecting analytical process quality, which constrains their potential.
    Approach: They propose a novel reward model that leverages self-reflection to assess analytical quality and enhance preference modeling.
    Outcome: The proposed model improves performance on four benchmarks and significantly mitigates positional bias.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations